[Feat] Served Laya as a decision model (--model laya-served) - #35
Open
cacheline999 wants to merge 14 commits into
Open
cacheline999 wants to merge 14 commits into
cacheline999 wants to merge 14 commits into
Conversation
laya 0.3.9 builds the encoder under transformers' no_init_weights inside laya.load and loads the checkpoint with strict=True, so the without_weight_init() wrapper from ThinkFlowLab#22 and its test stubs are no longer needed. The lock moves laya from 0.3.5 to 0.3.9, the first release with the skip; later releases change MPS precision and are left for their own bump. The system_one config-error test pins the version lookup, so it passes with the laya extra installed as well as on the core install. Closes ThinkFlowLab#28
system1-omni's Laya worker pins laya[serve]==0.3.20. Comparing in-process Laya with the served one (ThinkFlowLab#20) needs the same library on both sides, so the in-process extra moves from 0.3.9 to 0.3.20. 0.3.20 autocasts to fp16 on MPS for requests with five or more questions; CPU runs and smaller requests are unchanged.
docs/served-laya.md designs the served Laya decision model: requirements, the target interface, what today's server lacks and where each part gets built, the client (--model laya-served), error handling and trade-offs. docs/api/laya-systemone.openapi.yaml is the target interface as OpenAPI 3.1 (RFC 9457 errors with stable codes, request ids, Server-Timing, /livez and /readyz, 503 with Retry-After, served_by per response); each item is marked implemented or planned. laya-systemone.current.openapi.yaml specifies today's server (laya-serve 0.3.20 behind the system1-omni worker and frontend), checked against traffic from a running worker.
ServedLayaModel asks a Laya served over HTTP (system1-omni's worker, its omni-jev frontend, or plain laya-serve) through POST /v1/systemone, with the body the in-process model builds (laya_question). The client holds one deadline per decision, retries once on a dropped connection, 502 or 504, and after Retry-After on 503, sends one X-Request-Id per decision, and maps every other status to MODEL_CALL_FAILED or MODEL_SERVICE_CONFIG_ERROR. Decision.model is the served checkpoint and revision, read from the response's served_by when a server sends it, else from /health (refreshed after 30 s, since the worker reports a CPU fallback there), else from routing.repo against plain laya-serve. The window check is shared with LayaModel (check_window). Tests run over httpx.MockTransport: the shared contract, the error and retry mapping, identity, and responses recorded from a real worker, the frontend and laya-serve, which are checked against the as-implemented OpenAPI spec (jsonschema, pyyaml and referencing join the dev extra).
--model laya-served on every agent, on decide, on the rails and in MCP decide. A test finds any name list, match or Literal that has laya without laya-served, so a new front cannot miss it.
…nkFlowLab#20) A tool-front tick keeps Decision.model as model and, from a served model, served_by, url, request_id and server_timing (Decision.provenance), so a run's artifacts show the checkpoint, revision and device of each step.
decision-models.md and configuration.md list laya-served and its variables; served-laya.md gains a run section (start the worker once, use it from the CLI and MCP) and matches system1-omni#30 on /health freshness and GPU-only worker options.
A refused connection after the retry now says the worker listens only once warm and points to system1-omni's recipe; a first /health read that fails no longer logs about a previous reading that does not exist. The --model help of every front names laya-served.
90 decisions per configuration on an M1 Pro: in-process Laya, the system1-omni worker (compile + fp16) and the same worker behind omni-jev routed every (seed, ticket) pair the same and scored 63/90 each; p50 101, 77 and 82 ms. compare_served.py builds the table from the job dirs and keeps every decision in served_laya_records.json.
The worker now starts as `python -m frontend.laya_mps --compile --weights fp16` and its /health reports `compile.enabled` instead of `compile.mode`. served_by records `compiled` (true/false); both specs, the run section and the not-up message follow. Worker fixtures were recorded again at 3d6cb57; the other responses came out byte for byte the same. The ticket-router comparison was rerun at 3d6cb57: 63/90 in every configuration, all 90 decisions routed the same, p50 88 ms in process and 72 ms served, direct and through omni-jev.
…FlowLab#20) The registration check read s1a's sources with the platform encoding, cp1252 on Windows, and reported paths with backslashes there. Every new text read and write now names utf-8, and paths are reported in POSIX form.
…stopped episodes (ThinkFlowLab#20) The identity refresh every 30 s ran before the decision's deadline began, so a stalled /health could stretch one decision to LAYA_SERVED_TIMEOUT_S plus 2 s. The refresh now shares the decision's deadline and takes at most half of what is left. compare_served.py paired the batch's ticket ids with the ticks, so a job with an episode that stopped early raised in zip(). It now pairs the processed routes with the ticks and checks that each pair agrees.
…rors (ThinkFlowLab#20) The not-up error quoted the M1 Pro ready time with --compile and fp16, which misleads on a CPU worker or another Mac; it now says the worker may still be starting and points to the recipe. The timeout comment gives its reason instead of a measured latency range, and the browser comment names served Laya as the HTTP one. The run section quotes the real message.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Closes #20.
--model layaloads Laya in every process, so each CLI run and each MCP call pays the load and nothing shares a warm model. system1-omni can now keep Laya warm in a worker (ThinkFlowLab/system1-omni#30), and this PR lets agents use it.How
s1a/decision_models/served.pyis the new backend. Open it first.ServedLayaClientposts to/v1/systemoneand owns the connection handling:LAYA_SERVED_TIMEOUT_S, default 5 s);Retry-Afteron a 503;X-Request-Idper decision, reused on the retry;MODEL_CALL_FAILEDorMODEL_SERVICE_CONFIG_ERROR.ServedLayaModelbuilds the same body as in-process Laya withlaya_question(). It shares the full-window check withLayaModel, nowcheck_window()inlaya.py.It also records who answered:
Decision.modelis<checkpoint>@<revision>. The details come from the response'sserved_bywhen a server sends one. Otherwise they come from the worker's/health, re-read after 30 s, since the worker reports there when Laya falls back to the CPU. Plain laya-serve reports less, so the record keepsrouting.repofor it.The contract is written down in
docs/served-laya.md, as design and trade-offs. There are two OpenAPI 3.1 specs indocs/api/:laya-systemone.current.openapi.yamlis today's servers, checked against recorded traffic;laya-systemone.openapi.yamlis the target, with problem+json errors,served_by,Server-Timingand/readyz.The gap between the two belongs to system1-omni; section 4 of the doc lists where each part would go.
What
--model laya-servedworks on every agent, ondecide, on the rails and in MCPdecide.tests/test_served_laya_registration.pyfails if a new list offerslayawithout it.LAYA_SERVED_URLis required: the worker,omni-jevor laya-serve.LAYA_SERVED_MODEL,LAYA_SERVED_API_KEY,LAYA_SERVED_TIMEOUT_S,LAYA_SERVED_MAX_LEN.modelfor every decision model. A served model's ticks also keepserved_by,url,request_idandserver_timing.decision-models.md,configuration.md,.env.example, the CHANGELOG, and a run section inserved-laya.md.jsonschema,pyyamlandreferencingfor the spec check (pyproject.toml,uv.lock, CONTRIBUTING). They were already in the environment throughmcpandopenjiuwen.layaextra to 0.3.20, the version the worker runs.The worker features used here (warm before listening, a live
/healthwith checkpoint, revision and device) come from ThinkFlowLab/system1-omni#30, which is still open. Against plain laya-serve everything works, with the checkpoint as the only identity. The run section links the recipe on the PR branch until it merges.Results
ticket_router on an M1 Pro, 3 seeds × 30 tickets per configuration. Laya is checkpoint
55cf4c4with laya 0.3.20. The worker is system1-omni3d6cb57(the head of ThinkFlowLab/system1-omni#30), started with--compile --weights fp16. Details, commands and every decision are inevals/ticket_router/SERVED_LAYA.md.--model laya, MPS, fp32)omni-jevAll three routed every one of the 90 decisions the same. In-process and worker probabilities differ by at most 0.001. In-process Laya is slower here because it runs fp32 and uncompiled, not because of HTTP.
Open questions
laya-served, chosen so records tell served runs from in-process ones.docs/api/laya-systemone.openapi.yaml. If it looks right, the worker-side parts belong in system1-omni.Verification
uv run ruff format --check . && uv run ruff check . && uv run ty check: all pass.uv run pytest -qandscripts/smoke.sh: 586 passed and 43 skipped with the laya extra installed;smoke: ok. With torch, laya and transformers blocked from import, the new tests pass (113 passed, 2 skipped), as on CI's core install.CHANGELOG.mdand the docs say what the code does now.New tests run over
httpx.MockTransport:DecisionModelContract;tests/data/served_laya/holds 27 responses recorded from the worker, fromomni-jev(with the worker up and down) and from laya-serve. They validate against the as-is spec.Live on the same M1 Pro and worker:
decidevia the CLI and over MCP (stdio)injection_guardrailomni-jevomni-jev--device cpu) withLAYA_API_KEYsetLAYA_SERVED_API_KEY; with it, 21/30 and ticks showcpu,float32, not compiled